Skip to content

fix: remove the per-call struct format strings from the response parse path - #286

Merged
jaysonsantos merged 2 commits into
mainfrom
perf/response-struct-unpack
Sep 19, 2026
Merged

jaysonsantos merged 2 commits into
mainfrom
perf/response-struct-unpack

Conversation

@jaysonsantos

Copy link
Copy Markdown
Owner

Base branch: modernize/python-310-floor (#278). Review that pull request
first. The diff here is incremental, and it is nine lines.

What changed

bmemcached/protocol.py gains a module-level FLAGS_UNPACKER = struct.Struct('!L').
get and get_multi use it to read the 4 byte flags field, then slice the
rest of the body. Neither builds a format string any more.

The bytes() copy in _read_socket stays. The reason is below, with numbers.

Why

Commit 72d2aaf precompiled the request packers, but that work covered the send
path only. The receive path still built a format string on every call:

struct.unpack('!L%ds' % (bodylen - 4), ...)                      # get
struct.unpack('!L%ds%ds' % (keylen, bodylen - keylen - 4), ...)  # get_multi

CPython caches compiled struct formats in an LRU cache, 100 entries by
default. Because these formats embed the per-call value length, a workload
with more than 100 distinct value sizes evicts entries and recompiles on every
call. This is the same pattern 72d2aaf removed on the send side.

A struct.Struct for '!L%ds' % n is specific to one length n, so
precompiling one per length does not help. A fixed-format read plus a slice
compiles nothing at all.

Verification

nix develop --command bash -c 'pytest -q'

Result: 261 passed. This matches the baseline on main.

nix develop --command bash -c 'flake8'

Result: 0 errors.

Benchmark

The benchmark runs the parse step only. No socket and no server are involved,
because the change is about format compilation. Python 3.12, timeit.

Case Before After Speedup
get, 500 distinct value lengths (past the LRU cache) 0.418 us 0.137 us 3.05x
get, 20 distinct value lengths (inside the cache) 0.227 us 0.129 us 1.77x
get, one repeated length (best case for the cache) 0.227 us 0.121 us 1.87x
get_multi, 500 distinct value lengths 0.495 us 0.187 us 2.65x
get_multi, one repeated length 0.292 us 0.173 us 1.69x

The gain is largest past the cache, as expected. It is still real inside the
cache, because a % format and a lookup cost more than a slice.

The _read_socket copy: measured, and kept

The issue asked for a measurement, not a theory. Here it is.

The copy is real. bytes(bytearray) costs 0.073 us at 64 bytes, 0.111 us at
1 KB, 1.80 us at 16 KB, and 17.4 us at 256 KB. At large value sizes it
dominates the parse step this pull request just made faster.

It still stays, because removing it changes public return types. A bytearray
slice is a bytearray, not bytes:

  • deserialize returns the raw buffer for a value carrying the binary flag.
    get would return a bytearray in place of bytes.
  • get_multi uses the response key as a dict key. A bytearray is
    unhashable, so this raises TypeError.
  • stats() tests isinstance(key, bytes) and would stop decoding the key. It
    would then use an unhashable bytearray as a dict key.
  • The error paths format extra_content into an exception message. The text
    changes from b'...' to bytearray(b'...').

Each of those needs its own bytes() call. That puts the copy back for the
binary path and widens the change well past this issue. The acceptance
criteria allow the copy to stay with a stated reason, so it stays.

The copy is worth its own issue, together with a decision about whether
returning bytes for binary values is a contract this project wants to keep.
I have not opened that issue, because the answer is a design call for you, not
a defect. Say the word and I will open it.

Risks

A reviewer must check two points.

  1. The slice arithmetic must match the old format strings. For get, the
    old format read 4 bytes of flags then bodylen - 4 bytes of value, so the
    value is extra_content[4:]. For get_multi, the old format read 4 bytes
    of flags, then keylen bytes of key, then bodylen - keylen - 4 bytes of
    value, so the key is extra_content[4:4 + keylen] and the value is
    extra_content[4 + keylen:]. extra_content is exactly bodylen bytes
    long, which is what makes the open-ended slices correct.
  2. Return types are unchanged. _read_socket still returns bytes, so
    every slice is bytes, exactly as struct.unpack produced before. This is
    the direct consequence of keeping the copy.

struct.unpack raised struct.error on a body whose length did not match the
header. A slice does not. A short body now yields a short value in place of an
exception. No test covered that path, and the header check that #273 adds
catches the desynchronized-stream case that produces it.

Closes #276

@coderabbitai

coderabbitai Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 39 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 2 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 3ccc506e-33b8-4261-a449-dd6b4afa5617

📥 Commits

Reviewing files that changed from the base of the PR and between 046f849 and fd8178d.

📒 Files selected for processing (1)
  • bmemcached/protocol.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-08T23:33:07.731187Z 7f3b519 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Base automatically changed from modernize/python-310-floor to main September 9, 2026 14:55
@cursor
cursor Bot force-pushed the perf/response-struct-unpack branch from 7f3b519 to 5e4c671 Compare September 19, 2026 19:50
@cursor

cursor Bot commented Sep 19, 2026

Copy link
Copy Markdown

Rebased onto current main (046f849). Tip is fcd5548.

Conflicts were only in bmemcached/protocol.py at the get / get_multi unpack sites. Kept this PR’s FLAGS_UNPACKER + slice path; left main’s later modernization (ruff, pyproject, conftest, client/protocol fixes) in place. The already-merged #278 commits were not replayed.

Follow-up style: commit matches main’s ruff quotes/slices. Local pytest -q: 271 passed. GitHub mergeable: MERGEABLE. Remaining BLOCKED is reviews / in-flight required checks, not conflicts. claude-review failed because the pusher is cursor[bot].

@mergify

mergify Bot commented Sep 19, 2026

Copy link
Copy Markdown

Tick the box to add this pull request to the merge queue (same as @mergifyio queue).

  • Queue this pull request

@jaysonsantos jaysonsantos changed the title perf: remove the per-call struct format strings from the response parse path fix: remove the per-call struct format strings from the response parse path Sep 19, 2026
jaysonsantos and others added 2 commits September 19, 2026 19:58
…e path

Commit 72d2aaf precompiled the request packers. That work covered the send
path only. The receive path still built a format string on every call.

get built '!L%ds' % (bodylen - 4). get_multi built
'!L%ds%ds' % (keylen, bodylen - keylen - 4). CPython caches compiled struct
formats in an LRU cache with 100 entries by default. A workload with more
than 100 distinct value lengths evicts entries and recompiles on every call.

A Struct for '!L%ds' is specific to one length, so precompiling one per
length does not help. Both sites now read the fixed 4 byte flags field with
a module-level struct.Struct('!L') and slice the rest. A slice compiles
nothing.

Measured on Python 3.12, parse step only, no socket:

  get, 500 distinct value lengths:        0.418 us -> 0.137 us  (3.05x)
  get, 20 distinct value lengths:         0.227 us -> 0.129 us  (1.77x)
  get, one repeated length:               0.227 us -> 0.121 us  (1.87x)
  get_multi, 500 distinct value lengths:  0.495 us -> 0.187 us  (2.65x)
  get_multi, one repeated length:         0.292 us -> 0.173 us  (1.69x)

The bytes() copy in _read_socket stays. The issue asked for a measurement
and a decision.

The copy is real. bytes(bytearray) costs 0.073 us at 64 bytes, 0.111 us at
1 KB, 1.80 us at 16 KB, and 17.4 us at 256 KB.

Removing it changes public return types, because a bytearray slice is a
bytearray:

- deserialize returns the raw buffer for a value with the binary flag. get
  would return a bytearray in place of bytes.
- get_multi uses the response key as a dict key. A bytearray is unhashable.
- stats() tests "isinstance(key, bytes)" and would stop decoding the key. It
  would then use an unhashable bytearray as a dict key.
- The error paths format extra_content into an exception message. The text
  changes from b'...' to bytearray(b'...').

Each of those needs its own bytes() call, which puts the copy back for the
binary path and widens the change well past this issue. The acceptance
criteria allow the copy to stay with a stated reason. This is the reason.

Refs #276
Match main's ruff quote and slice style on FLAGS_UNPACKER.

Co-authored-by: Jayson Reis <santosdosreis@gmail.com>
@cursor
cursor Bot force-pushed the perf/response-struct-unpack branch from fcd5548 to fd8178d Compare September 19, 2026 19:59
@cursor

cursor Bot commented Sep 19, 2026

Copy link
Copy Markdown

Commit type changed from perf: to fix: to match 72d2aaf.

@jaysonsantos
jaysonsantos merged commit d123255 into main Sep 19, 2026
9 of 10 checks passed
@jaysonsantos
jaysonsantos deleted the perf/response-struct-unpack branch September 19, 2026 20:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Remove per-call struct format strings from the response parse path

2 participants